Accessibility settings

Published on in Vol 28 (2026)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/91968, first published .
Four seniors reminisce over old photos at a table with flowers and tea.

Quality of Life of People Living With Dementia Residing in Nursing Homes: Secondary Analysis of Observational Data

Quality of Life of People Living With Dementia Residing in Nursing Homes: Secondary Analysis of Observational Data

1Alzheimer Centrum Limburg, Faculty of Health, Medicine and Life Sciences, Maastricht University, Doctor Tanslaan, 12, Maastricht, Limburg, The Netherlands

2Department of Health Service Research, CAPHRI Care and Public Health Research Institute, Faculty of Health Medicine and Life Science, Maastricht University, Maastricht, Limburg, The Netherlands

3The Living Lab in Ageing & Long-Term Care, Maastricht, The Netherlands

4Institute for Communication, Media and Information Technology, Rotterdam University of Applied Sciences, Rotterdam, The Netherlands

5Data Supported Healthcare, Research Center Innovations in Care, Rotterdam University of Applied Sciences, Rotterdam, The Netherlands

6Faculty of Medicine and Health Sciences, Macquarie University Sydney, Sydney, Australia

7Medical Delta Livinglab: Data Supported Healthcare & Innovation, Delft, The Netherlands

8School of Population Health, Royal College of Surgeons in Ireland, Dublin, Ireland

Corresponding Author:

Dirk Steijger, MSc


Background: Quality of life (QoL) plays a crucial role in dementia care; however, QoL and its dynamic, context-dependent nature can be difficult to capture among people living with dementia due to challenges in memory and communication, and limitations of self-reported QoL instruments. Observational tools such as the Maastricht Electronic Daily Life Observation (MEDLO) provide narrative descriptions of the daily life of people living with dementia in nursing homes. However, the MEDLO tool was not developed to assess QoL specifically, and it remains unclear to what extent its narrative descriptions reflect aspects of QoL. Analyzing these narrative descriptions is labor-intensive and time-consuming. Recent advances in natural language processing, including large language models (LLMs), offer the potential to analyze these narrative descriptions at scale.

Objective: The study aims to explore whether an LLM can be used to structure existing MEDLO narrative data into interpretable QoL-relevant patterns in people living with dementia. Specifically, this study examines whether N-gram analysis, sentiment analysis, and LLM-based topic modeling can identify recurring language patterns, emotional tone, and semantic clusters that can be mapped to Lawton QoL domains.

Methods: This study conducted a secondary analysis of existing MEDLO observational data from 151 people living with dementia residing in Dutch long-term care. Narrative data had been documented by trained observers, describing activities, interactions, settings, and emotional expressions. For analysis, a local secure pipeline was developed in which GPT-4o-mini was deployed. The pipeline comprised three analytical steps: (1) N-gram frequency analysis, (2) sentiment analysis, and (3) topic modeling. Prompts were iteratively refined through prompt engineering. Coauthors and domain experts reviewed outputs for coherence, contextual plausibility, and relevance to long-term care practice.

Results: A total of 5622 narratives (50,106 words) from 151 people living with dementia were analyzed. The narratives were short, averaging 10.5 (SD 5.80) words per narrative. N-gram frequency analysis identified the frequent documentation of passive activities (sits at the table) in limited indoor settings (living room). Emotional well-being was often described in positive terms (smiles and laughs), whereas explicitly negative expressions (cries and distress) occurred less frequently. Weighted sentiment analysis showed that, although fewer in number, negative expressions carried a stronger intensity, resulting in an overall predominance of negative sentiment across all QoL domains. Topic modeling identified 8 coherent clusters, most of which mapped onto multiple QoL domains, underscoring QoL’s multidimensionality.

Conclusions: LLM-based analyses identified predominantly passive activities with little variation in indoor settings, while people living with dementia were often described as having positive affect. This exploratory study suggests that LLM-based analyses may help structure observational narratives into QoL-relevant patterns, but further validation is needed before such outputs can inform person-centered care practice.

J Med Internet Res 2026;28:e91968

doi:10.2196/91968

Keywords



Dementia is a neurocognitive disorder with cognitive decline, leading to impairments in cognitive, social, and behavioral functioning [1,2]. While the symptoms of dementia can negatively affect quality of life (QoL), factors such as individual needs, lived experiences, and environmental support also play a crucial role [3-5]. Lawton conceptualizes QoL in people living with dementia as a multidimensional concept, consisting of 4 domains: behavioral competence (physical health and functional ability), psychological well-being (mood), objective environment (environmental factors), and perceived QoL (the individual’s own evaluation of their functioning, mood, and overall life) [3,6]. This holistic perspective emphasizes that QoL extends beyond symptom severity, with consideration of the way individuals interact with their environment and preserve a sense of identity and autonomy. Improved QoL in people living with dementia has been linked to reduced agitation and depression and to an improved sense of autonomy and satisfaction with care [4,7-10].

As dementia progresses, declines in communication, memory, and executive functions often make it difficult for people living with dementia to recall and express their preferences, emotions, and sources of discomfort. This can reduce the usefulness of self-report QoL measures [11,12]. Other measures include proxy-based questionnaires and systematic observational methods. However, these methods are not always capable of reflecting the dynamics and nuances associated with day-to-day QoL [13,14]. In recent years, observation tools such as the Maastricht Electronic Daily Life Observation (MEDLO) have introduced a more comprehensive assessment approach by integrating quantitative information with narrative descriptions of the behavior, mood, and environment of people living with dementia in nursing homes [15]. These narrative descriptions, hereafter referred to as narrative data, are collected by a trained observer and contain written information about the behavior of an individual and the environmental contextual information in which this behavior occurred. They capture detailed, real-time descriptions of the daily lives of people living with dementia in long-term care settings, offering a unique window into the everyday QoL. However, the MEDLO tool was not developed to assess QoL specifically, as it does not include a predefined QoL scoring system, validated QoL subscales, or direct self-reported evaluations of QoL by the person living with dementia. Therefore, the narrative data may contain signals that are relevant to QoL, but these signals remain implicit. Moreover, manual analysis of these narrative data is time-consuming, difficult to scale, and susceptible to interpretative bias [16-20]. Consequently, large volumes of narrative data collected with the MEDLO tool remain underanalyzed and underutilized for deriving QoL-related insights in dementia care.

However, recent advances in AI, particularly in natural language processing (NLP), hold at least a partial solution to the practical limitations of current methods for analyzing narrative data in long-term care. NLP focuses on the analysis and interpretation of human language in textual form by enabling automated computerized detection of themes, tones, and meaning within large volumes of narrative data [21]. In recent years, large language models (LLMs) have represented a significant advance in NLP, enabling models to interpret linguistic patterns with greater contextual and semantic depth than previous alternative approaches [22,23]. Trained on vast corpora of text, LLMs are capable of analyzing large volumes of text that would otherwise be impractical to analyze due to limited resources. LLM-based analysis can systematically extract recurring language patterns, emotional tone, and semantically related themes from large volumes of unstructured narrative data [24,25]. Therefore, LLMs could potentially also be used to identify QoL-relevant patterns in unstructured narrative data, with the potential to transform these data into interpretable QoL-related information [26]. Recent studies have used NLP and LLMs to extract information from free-text electronic health records [27-32]. Few studies have applied these methods within dementia care [27,33-35], and none have focused on QoL.

The study aims to explore whether LLM-based analyses can be used to structure existing narrative data, stemming from MEDLO data, into interpretable QoL-relevant patterns among people living with dementia residing in nursing homes in the Netherlands. Specifically, we examine whether N-gram analysis, sentiment analysis, and LLM-based topic modeling can identify recurring language patterns, emotional tone, and semantic clusters that can be mapped to Lawton QoL domains. As MEDLO was not developed as a QoL instrument, these outputs are interpreted as exploratory QoL-relevant signals rather than direct measurements of QoL. This study contributes to the field by showing how routinely collected observational text may be transformed into structured, QoL-relevant patterns for dementia care research


Study Design

This study builds on a previous exploratory observational study by analyzing the existing narrative data to explore whether LLM-based analyses can transform these narratives of daily life into interpretable insights of QoL in people living with dementia.

Data Collection

The dataset used in this study was originally collected as part of the “Green Care Farms—an innovative care environment for adults living with dementia” project [36]. Similar to the current study, the Green Care Farms project was conducted within the Living Lab in Ageing and Long-Term Care in South Limburg, the Netherlands [37]. The Green Care Farms project aimed to monitor the daily lives of people living with dementia residing in Dutch care farms and traditional nursing homes, examining how the care farm environment influences their well-being.

Data were collected using the MEDLO tool, a tablet-based observation instrument specifically developed to capture detailed, real-time descriptions of the daily lives of people living with dementia in long-term care settings [14,15].

Data were collected over a period of 2 months per Green Care Farm, and the period of the day of the observation was selected randomly. Three researchers and 2 research assistants carried out observations at the 4 green care farms. All observers were trained by an experienced user of the MEDLO tool. This training consisted of studying the manual, practicing the training, and discussing the data. Within this training, both researchers observed the same people living with dementia and discussed their findings. The researchers conducted observations independently when they reached a complete agreement.

People living with dementia were observed for one morning (7 AM-11:30 AM), one afternoon (11:30 AM-4 PM), and one evening (4 PM-8:30 PM), with a 30-minute break each day. Every 20 minutes, a maximum of 11 people living with dementia were observed for 1 minute in a random sequence. This led to 12 observations per individual living with dementia per observation day and 36 momentary assessments per individual living with dementia in total. The people living with dementia were observed for 1 minute. By randomly selecting time points within the days of a person with dementia, insight into their daily lives was gained [15]. All observed people living with dementia had a formal diagnosis of dementia, confirmed by medical staff.

During each 1-minute observation, observers recorded both quantitative scores and narrative descriptions on four domains of daily living: (1) activity performance, (2) physical environment, (3) social interaction, and (4) emotional well-being. The narrative descriptions were written as free-text entries in natural language, without predefined categories or templates, allowing observers to capture contextual details and subtle nuances often overlooked in standardized formats [13,14]. For example (translated from Dutch): “She is sitting at the dining table next to the kitchen unit, staring at a newspaper in front of her.” These narrative descriptions served as the primary dataset for the current study’s NLP analysis.

Ethical Considerations

Data were derived from the Green Care Farms study. This study was reviewed and approved by the Medical Ethics Committee of Zuyderland, the Netherlands (METCZ20210097). An amendment to the study protocol received a positive decision from the same ethics committee on February 20, 2024. The present study used anonymized observational data collected within this approved study protocol for secondary analysis. All procedures were conducted in accordance with the ethical standards of the institutional and/or national research committee and with the Declaration of Helsinki and its later amendments [38]. The data used in the present study had been anonymized by the researchers from the original study before being transferred to the current research team.

Data Preprocessing

Preprocessing is a critical step in text mining, as it prepares raw text for analysis by removing noisy or unreliable data that may interfere with accurate interpretation [19,39]. In this context, “noise” refers to inconsistencies in the data that do not carry meaningful information for analysis [40]. For the current study, preprocessing and analysis were performed using Python version 3.11.8, a free software package for general-purpose programming language [41]. The code for all the analyses can be found on GitHub [42].

While traditional NLP pipelines often involve extensive text cleaning, such as tokenization, normalization, stop-word removal, stemming, and handling punctuation and numbers, minimal preprocessing was conducted in this study [43]. The rationale for this approach was to preserve the natural language, contextual detail, and subtle nuances in the narrative data [44]. These nuances, including emotional tone and subtle behavioral cues, are essential for understanding the QoL of people living with dementia.

Although this study applied minimal preprocessing, several preprocessing steps were carried out in collaboration with researchers from the original Green Care Farm project, leveraging their familiarity with the data to ensure high data quality. (1) All narrative data from the MEDLO tool were extracted from the Excel file into a plain text file and loaded into Python; (2) all words were lowercased to enable consistent matching of identical terms regardless of case and to facilitate abbreviation expansion; (3) commonly used abbreviations, introduced during data collection to expedite note-taking, were replaced with their full-length equivalents (eg, lr: living room, bk: big kitchen); (4) double spaces and stray blank spaces were removed, and numbers were written out in words to improve textual consistency; and (5) the final preprocessed text file was converted into a PDF format, which was then used as input for LLM-based analyses, allowing the model to process the narrative data in its original form. The conversion to PDF was a pragmatic workflow decision rather than an analytical preprocessing step. The secure analysis pipeline had originally been developed to process PDF input files, and therefore, the final preprocessed text file was converted to PDF to ensure compatibility with this existing pipeline.

Model Architecture

This study analyzed the narrative data using an LLM approach, using GPT-4.0-mini deployed through the Azure OpenAI Service (Microsoft) [45]. The analysis was conducted in Python 3.11.8 using the OpenAI Python package (openai==1.59.6), which was used to send requests to the Azure OpenAI chat-completion end point [46,47]. Supporting packages included PyPDF2 for extracting text from the PDF input file and pandas for structuring the extracted text during analysis [48,49]. The extracted narrative text was combined into a single Dutch text corpus and submitted to the deployed model through the local Azure deployment named 4OMINI using API version 2023-09-01-preview. The request specified a maximum generated output length of 4096 tokens. Temperature, seed, and structured response format were not explicitly set in the request. Therefore, outputs were generated without explicitly controlling randomness or enforcing a structured response format [50-52]. The analyses were performed via the model between November 2024 and May 2025. The model was not trained or fine-tuned on these data but was used only for inference. All preprocessing, prompt preparation, and output storage were conducted within the local secure research environment of Maastricht University. This secure and scalable blueprint of the model enables organizations to use LLMs, such as GPT-4.0-mini, while safeguarding data privacy.

Prompt Engineering

Overview

Prompt engineering was used to instruct the model in extracting language pattern insights from the narrative data. Three task-specific prompts were developed in Dutch: (1) N-gram identification and QoL-domain mapping, (2) sentiment scoring of N-grams, and (3) topic modeling with QoL-domain mapping. Each prompt specified the analytical task, the expected output format, the use of Lawton QoL domains as the conceptual framework, and instructions to remain close to the wording of the original observational narratives. The final prompts can be found at the repository at GitHub [42].

Prompt development followed an iterative refinement process [53]. After each analytical run, coauthors (DS, MCS, RFvdW, SA) evaluated whether the output was coherent, sufficiently specific, grounded in the narrative data, and consistent with the predefined QoL-domain framework. Prompt adjustments were made when outputs were too broad, semantically unclear, or inconsistent with the predefined QoL-domain framework. This iterative process continued until the coauthors reached consensus that the outputs were sufficiently interpretable, contextually plausible, and aligned with the study aim. In total, the N-gram identification and QoL-domain mapping prompt were run 9 times, the sentiment scoring prompt was run 6 times, and the topic modeling and QoL-domain mapping prompt were run 10 times. All prompt versions and corresponding outputs were stored to document the development process and are available from the corresponding author upon request. The final prompts were each executed once for the final analysis and were not repeated multiple times as identical runs to formally assess run-to-run stability. Therefore, some degree of output variability cannot be ruled out.

All analyses were conducted on the original Dutch narrative data. The narratives were not translated before analysis. English translations of selected examples are provided only for reporting purposes in this paper.

N-Gram Frequency Analysis and Sentiment Analysis

The first analysis performed by the model was N-gram frequency analysis, and the second analysis was sentiment analysis. To accomplish both analyses, 2 different prompts were developed to guide the model’s analyses: one prompt to identify the most frequent N-grams per QoL domain and the other prompt to determine the most positive and most negative sentiment-scoring N-grams per QoL domain.

An N-gram is a sequence of N consecutive words extracted from a text [54]. For example, the sentence “The resident smiled while listening to music” can be broken down into the following: (1) unigrams: “The,” “resident,” “smiled,” “while,” “listening,” “to,” “music;” (2) bigrams: “The resident,” ‘‘resident smiled,” “smiled while,” “while listening;” and (3) trigrams: “The resident smiled,” “resident smiled while,” “smiled while listening.” Using the first prompt, the model was instructed to identify the N-grams across the narrative data. The next step in this prompt was to categorize each N-gram into an appropriate QoL domain. The mapping of N-grams to QoL domains was operationalized using the Lawton QoL framework. The definitions of behavioral competence, psychological well-being, objective environment, and perceived QoL were added to the context section of the prompts. Based on these definitions, the model was instructed to assign each N-gram to the most appropriate QoL domain. Consequently, the prompt instructed the model to report the 15 most frequent N-grams per QoL domain. The 15 N-grams assigned to each QoL domain were then reviewed by the coauthors and cross-referenced with the literature to assess whether they were contextually plausible and consistent with the domain to which they had been assigned.

In cases where an N-gram appeared too ambiguous to convey a clear meaning, the model was instructed, using the first prompt, to extend the N-gram by one additional word, derived from the original narrative context, rather than generated independently, in order to improve the clarity of the N-gram. For instance, the trigram “is happy with” could be extended to “is happy with visits” depending on its original context. This approach could ensure that N-grams reflect meaningful language units rather than isolated, ambiguous terms. Consequently, using the second prompt, the model identified the most positive and most negative sentiment-scoring N-grams categorized per QoL domain. Specifically, this prompt first identified all N-grams present in the narrative data and then applied sentiment analysis to all N-grams. Thus, sentiment analysis was not performed on the most frequent N-grams, from the first prompt, but on all N-grams identified in the narrative data. N-gram sentiment analysis involves determining the emotional tone of a given text [19,39,55]. For example, “The resident smiled while listening to music” would typically be classified as positive, whereas “The resident sat alone and looked down without responding” would be classified as negative. Neutral sentiment could be applied to emotionally neutral statements such as “The resident walked across the room.”

Each identified N-gram was assigned a sentiment score on a scale from −1 (“very negative”) to 1 or ambiguous sentences (“very positive”), with 0 indicating neutral sentiment. The model was instructed to apply both lexicon-based and context-aware sentiment score approaches to improve the accuracy of the sentiment scores assigned to the N-grams. Lexicon-based sentiment analysis relies on predefined emotional associations of individual words; for instance, “happy” is typically associated with positive sentiment. However, this method may misclassify more complex or ambiguous sentences. To address this complexity, a context-aware sentiment analysis was also used to interpret the meaning based on phrasing and tone. For example, although the sentence “The resident smiled briefly at the caregiver but avoided eye contact and stayed silent” included the word “smiled,” the overall sentiment might be neutral or negative due to the surrounding disengagement indicators.

While the most positive and most negative sentiment-scoring N-grams of each QoL domain were identified based on their sentiment scores, it should be noted that these N-grams might have appeared only once in the narrative data. As such, these most positive and most negative sentiment-scoring N-grams do not necessarily provide insight into a general sentiment tendency within each QoL domain. To gain insight into broader sentiment patterns, weighted sentiment summaries were calculated [56]. For each N-gram of the most positive and most negative sentiment-scoring N-grams in a QoL domain, the sentiment score of that N-gram was multiplied by its frequency of occurrence in the narrative data. Subsequently, the sum of these products was then divided by the total frequency of all the N-grams within each QoL domain. This yielded a weighted sentiment score for each QoL domain.

Topic Modeling With QoL Domain Mapping

The third analysis regarded topic modeling, which was used to identify recurring themes in the narrative data. This analytical task was instructed via a third prompt. Topic modeling refers to the process of grouping semantically related keywords into overarching themes [19,57,58]. For example, keywords such as pacing, fidgeting, frustration, wandering, and repetitive movement might be clustered under the topic label agitation and restlessness. In this study, topic modeling was implemented as an LLM-based analytical task rather than as a traditional statistical topic modeling [59,60]. The definitions of Lawton 4 QoL domains were included in the context section of the prompt. The model was instructed to group semantically related words into topics, assign a descriptive label to each topic, and then link each topic to one or more QoL domains (as a single topic might reflect more than one aspect of QoL) based on the associated words in the narrative data [60,61]. Coauthors then reviewed whether the topics with included keywords, their labels, and their assigned QoL domains were coherent, plausible, and consistent with the Lawton framework.

The prompt also included guidance to enhance analytical granularity, such as splitting broad topic clusters into more specific clusters. To improve the accuracy of the topic modeling analysis over time, the model was further guided, via the prompt, to suggest refinements to the prompt itself. As with the N-gram and sentiment analysis, prompt development followed an iterative process, with adjustments to structure, wording, and level of detail made in response to the relevance and coherence of earlier outputs.

Expert Model Evaluation

Expert evaluation was used to assess whether selected model outputs were meaningful and contextually plausible in relation to long-term care practice. For sentiment analysis, 2 researchers (LF and KR, both with 5 years of research experience in long-term care) from the original Green Care Farm project independently reviewed the top positive and negative N-grams per QoL domain [62]. They assessed whether the assigned sentiment direction and intensity were plausible, given the wording of the N-gram and its interpretation within the context of observational dementia care. Disagreements between reviewers were recorded and discussed, and the final interpretation was based on consensus.

For topic modeling, the model-generated clusters were reviewed by 4 members of the research team (DS, HC, SA, and MCS), with expertise in dementia care and QoL research, long-term care research, and health care AI methods. The topic clusters were assessed using 2 qualitative criteria: coherence and distinctiveness. Coherence referred to the internal consistency of each topic cluster, that is, whether the terms grouped under a given label were meaningfully related [58]. For example, a cluster containing the terms music, dancing, laughter, and clapping was assessed to be coherent and appropriately labeled as leisure activities. In contrast, if an unrelated term, such as agitation, appeared within that same cluster, it indicated a lack of coherence, which resulted in further fine-tuning of the prompt. Distinctiveness was assessed by examining the boundaries between topic clusters to avoid redundancy or overlap. If an overlap between topic clusters was identified by the research team, the model was iteratively prompted to generate more distinct clusters. This process of prompt engineering continued until the research team reached consensus that the topic clusters were sufficiently distinct, with minimal thematic overlap.


Overview

In total, 5652 observations were conducted on 151 people living with dementia during the Green Care Farms study [62]. Narrative descriptions were missing in 30 (0.53%) observations. The narrative descriptions totaled 50,106 words. An average of 10.5 (SD 5.8) words was noted per narrative. Table 1 presents the background characteristics of the 151 individuals living with dementia who were observed.

Table 1. Characteristics of people living with dementia observed in the Green Care Farm study.
CharacteristicsParticipants (N=151)
Age, mean (SD)84.7 (7.1)
Sex (female), n (%)108 (72)
S-MMSEa score, mean (SD)10.5 (7.4)
Barthel Index scoreb, mean (SD)13 (5.6)

aS-MMSE: Standardized Mini-Mental State Examination. The S-MMSE assesses global cognitive functioning, with lower scores indicating greater cognitive impairment.

bThe Barthel Index measures independence in activities of daily living, with lower scores indicating greater dependency. The table is adapted from the original study [62].

N-Gram Frequency Analysis

The N-gram frequency analysis revealed language patterns related to QoL. Frequently observed terms included laughing, smiling, satisfied, and nice, suggesting that people living with dementia were often described as being in a state of positive emotional well-being. Negative affective expressions, such as becoming angry, crying, and feeling distressed, appeared far less frequently than positive affective expressions.

Ten of the 15 N-grams within the QoL domain behavior competence reflected low-activity behavior. Observers frequently noted people living with dementia sitting at the table, reading the newspaper, or looking outside. Although terms that encompass “activity” such as helps with cooking or participates in gymnastics were present, they were less common. This suggests a pattern of predominantly passive engagement in daily life, despite generally positive emotional expressions.

The daily lives of people living with dementia were most often observed in 2 core settings: the private room and the living room. Additionally, a few observations of people living with dementia were made in outdoor spaces, such as a terrace, the garden, or in the sun.

Similar to other QoL domains, N-grams within the perceived QoL domain were mainly positive affective expressions, such as satisfied, nice, beautiful, and happy. The partial overlap in N-grams from the domain of psychological well-being was observed (ie, satisfied and happy). The 15 most frequent N-grams identified within each QoL domain are presented in Figure 1.

Figure 1 shows the 15 most frequent N-grams identified in the observational field notes for each of Lawton QoL domains. Frequencies indicate the number of times each N-gram occurred in the contextualized text segments included in the N-gram analysis. The panels show that psychological well-being was mainly represented by affective expressions such as “lachen,” “glimlacht,” and “blij;” behavioral competence by passive activity phrases such as “zit aan tafel” and “kijkt naar buiten;” objective environment by indoor settings such as “eigen kamer” and “woonkamer;” and perceived QoL by evaluative or experiential expressions such as “tevreden,” “leuk,” and “mooi.”

Although the model was instructed to extend ambiguous N-grams by one word to improve clarity, some N-grams still lacked semantic clarity and contextual relevance. For instance, certain N-grams live or direct reaction offered limited insight.

Figure 1. Most frequent N-grams in observational field notes, categorized per the Lawton quality of life (QoL) domain. (A) N-gram frequency psychological well-being domain; (B) N-gram frequency behavioral competence domain; (C) N-gram frequency objective environment domain; (D) N-gram frequency perceived QoL domain.

Sentiment Analysis

Figure 2 presents the 5 most positively and negatively scored N-grams within each QoL domain, showing a wide range from strongly negative to strongly positive sentiment. The analysis of the N-grams identifies distinct sentiment patterns across QoL domains. High positive sentiment-scoring N-grams (around +0.8 to+0.9) show engagement and enjoyment of people living with dementia, such as experiencing fun during games, being cheerful with visitors, and going walking with a group. On the contrary, strongly negative sentiment-scoring N-grams (around −0.8) appear in expressions such as crying, appearing confused, and feeling anxious, which predominantly capture emotional distress in people living with dementia.

Within each QoL domain, the sentiment polarity of the N-grams is clear. Perceived QoL ranges from female client seems satisfied to feels bad. Psychological well-being ranges from happy with visitors to crying. Behavioral competence ranges from goes walking with a group to complaining.

While these examples reflect a mix of positive and negative-scoring N-grams, the overall sentiment distribution in the narrative data was skewed. Across all QoL domains, negative sentiment outweighed positive sentiment. Positive sentiment made up between 17.9% and 24.9% of the total weighted sentiment, depending on the QoL domain. This suggests that the average sentiment per QoL domain was driven downward by the higher intensity of negative sentiment-scoring N-grams. Even when positive sentiment-scoring N-grams appeared more frequently, their lower sentiment scores were insufficient to offset the impact of fewer but more intensely negative sentiment-scoring N-grams. Table 2 presents the weighted positive and negative sentiment scores for each QoL domain.

Figure 2 shows sentiment scores for the top 5 most positive and negative N-grams categorized by Lawton QoL domains. Bars to the right of 0 represent positively scored N-grams, whereas bars to the left of 0 represent negatively scored N-grams. Colors indicate the QoL domain to which each N-gram was assigned: behavioral competence, psychological well-being, objective environment, or perceived QoL. These sentiment scores reflect the polarity of documented textual expressions and should not be interpreted as direct measures of participants’ QoL.

Figure 2. Sentiment scores of positive and negative N-grams by Lawton quality-of-life (QoL) domains.
Table 2. Weighted sentiment scores for the most positive and negative N-grams by Lawton quality-of-life (QoL) domains.
QoL domainPositive (wsuma)Negative (wsum)Pos/neg ratiob%Positivec%Negative
Behavioral competence43.1140.00.324.975.1
Psychological well-being39.9182.50.217.982.1
Objective environment38.9153.70.2520.279.8
Perceived QoL42.8135.00.324.475.6

awsum: weighted sum of sentiment scores, calculated as N-gram frequency multiplied by the sentiment score.

bPos/neg ratio: ratio of positive to negative weighted sentiment.

cPercentage weighted sentiment: proportion of positive or negative sentiment relative to the total weighted sentiment per QoL domain. These scores reflect sentiment patterns in documented observational text and should not be interpreted as direct measures of participants’ QoL.

Topic Modeling

Topic modeling identified 8 main topics in the narrative data. Each topic was defined by a set of representative keywords and was linked to one or more QoL domains, as shown in Table 3.

The model inferred thematic correlations based on the co-occurrence of keywords across different QoL domains. This showed that narrative data involving social interaction and daily activities (ie, eating and walking) frequently appeared alongside positive emotional expressions (ie, smiling and feeling relaxed). On the contrary, negative emotional states (ie, restlessness or sadness) were often reported in contexts of social isolation or reduced stimulation, such as when people living with dementia remained in bed or stayed alone in their rooms.

Perceived QoL was linked to the highest number of topics (6/8). The other QoL domains appeared in 4 topics. Many topics were assigned to multiple QoL domains, indicating that a single thematic cluster often reflects multiple QoL domains. For example, the topic eating and drinking was linked to behavioral competence, objective environment, and perceived QoL, reflecting the multidimensional nature of QoL.

Table 3. Overview of identified topics including keywords and associated Lawton quality-of-life (QoL) domains.
TopicaKeywordsQoL domains
Daily activitiesSleeping, eating, singing, reading, and cookingBehavioral competence, objective environment, and perceived QoL
Psychological well-beingLaughing, smiling, angry, confused, and contentPsychological well-being and perceived QoL
Communication with staff and coresidentsTalking, conversations, and singingPsychological well-being and perceived QoL
Physical activity and mobilityWalking, grocery shopping, walking around, and movingBehavioral competence and objective environment
Personal care and supportDressing, showering, medication, and helpingBehavioral competence and psychological well-being
Social interactionVisits, family, residents, and being togetherPsychological well-being and perceived QoL
Eating and drinkingSandwich, coffee, soup, tea, cookie, and mealBehavioral competence, objective environment, and perceived QoL
Environmental factorsTerrace, garden, kitchen, living room, sun, and flowersObjective environment and perceived QoL

aTopics were identified from the observational field notes and linked to Lawton QoL domains based on the content of the associated keywords.


Principal Findings

The study aimed to explore whether an LLM could be used to structure existing MEDLO narrative data into interpretable QoL-relevant patterns in people living with dementia residing in nursing homes in the Netherlands. Specifically, it examined whether N-gram analysis, sentiment analysis, and LLM-based topic modeling could identify recurring language patterns, emotional tone, and semantic clusters that could be mapped to the Lawton QoL domains.

A first implication from this study concerns the need to incorporate context into observational data when reporting and analyzing affect in people living with dementia. The findings from this study indicated that while people living with dementia were predominantly observed during passive activities and in a limited range of indoor settings, they were often observed expressing positive emotions during the day. In addition, the sentiment analysis showed that although negative expressions were documented less frequently, they received more extreme negative sentiment scores. This finding is consistent with the literature, suggesting that people living with dementia are frequently documented as comfortable and content in their daily lives; however, their well-being remains sensitive to moments when physical, emotional, or social needs are not fully met [3,63-65]. The reported high-intensity negative expressions may reflect moments when these needs are not adequately addressed [66,67]. However, measuring emotions in people living with dementia is challenging because their emotional expression is often characterized by neutral affect [63]. Consequently, in this study, observations reflecting neutral affect may have contributed little to the weighted sentiment scores, while less common negative expressions could have stood out against this neutral baseline and received relatively high-intensity scores, thereby influencing the weighted sentiment results.

Hence, there is a need to develop documentation guidelines to encourage observers to standardize their data collection methods and practices. While traditional NLP methods, such as term frequency-inverse document frequency, rely on word occurrences, LLMs interpret meaning from the surrounding context [23]. Consequently, the descriptive detail of the narrative data directly affects the outputs. This dependency on context-rich narrative data means that variation in how observers phrase their descriptions can lead to systematic differences in model outputs, even when experiences are similar [67-69]. For example, what to record, how detailed to be, and which emotions or behaviors to emphasize are each influenced by time constraints and professional judgment. Therefore, developing documentation guidelines could encourage observers to standardize their documentation. An example of such documentation guidelines is to include narrative information about the social context, setting, and activity, alongside emotional expressions, in care descriptions, which could provide richer input for LLM-based models while remaining feasible within care routines [70-72]. This also aligns with broader debates in health care–based LLM use, where the quality and style of input text have been shown to shape model performance in long-term care, potentially improving care for people living with dementia [70-72].

A second implication of this study concerns the ways in which LLM-generated conceptualizations of QoL can provide insight into how different aspects of QoL can co-occur and interact in dementia care and in what circumstances they do so [62,66,73-75]. The Kitwood personhood framework, along with later work on person-centered care, emphasizes that QoL derives from opportunities for purposeful activity, social reciprocity, and recognition of individuality rather than from momentary affective expressions alone [67,76,77]. From this perspective, positive documented expressions during passive activities in low-variation indoor settings may be interpreted as reflecting momentary positive affect, while leaving unanswered questions about the extent of meaningful engagement [78,79]. In our analysis, the LLM-generated outputs linked affective terms with descriptions of activity, agency, and social interaction. These textual patterns may provide contextual information that could support further exploration of the distinction between momentary positive affect and meaningful engagement. Furthermore, the topic modeling in this study provides support for the idea that all of Lawton QoL domains are interwoven [3,6,80]. For example, the description of enjoying helping with cooking may simultaneously reflect behavioral competence, psychological well-being, and perceived QoL. In the future, LLM-generated clusters may help describe how emotional, physical, social, and environmental aspects co-occur within observational narratives, while existing questionnaires currently isolate experiences into predefined categories [81-83].

In sum, this study aligns with broader developments in health care, where increasing attention is being given to the analysis of narrative data [24,84-87]. The current findings demonstrate that even brief narrative care data can be structured into recurring textual patterns. This creates opportunities to integrate near–real-time feedback loops into tools such as the MEDLO. Embedding LLM-based analysis within such tools may enable the immediate analysis of narrative descriptions, transforming routine documentation into dynamic QoL feedback of people living with dementia. For example, the automatic detection of subtle changes in emotional tone or activity diversity could help care teams recognize early signs of reduced well-being. Such insights may enable staff to identify emerging needs earlier and adjust activities, environments, or interactions in a more proactive and person-centered manner [67,88,89]. Importantly, these integrated feedback loops should complement, not replace, the expertise of care professionals by offering timely, data-informed insights to support daily practice [90,91]. Evidence from long-term care indicates that these technologies must be developed in close collaboration with domain experts, including care professionals, researchers, and technology developers, to ensure integration into daily workflows, meet usability requirements, and uphold ethical standards [92-95].

Strengths and Limitations

A key strength of this study was the unique nature of this real-world dementia-specific dataset. The manual analysis of this dataset would be very time-consuming. However, analyzing these data with an LLM enabled the systematic exploration of thousands of brief descriptions. Another strength lies in the decision to use minimal preprocessing, preserving the authenticity and contextual richness of the observational language. This approach enabled the LLM to capture the nuances of natural expression, thereby maintaining practical relevance in real-world care contexts. Transparency and reproducibility were also prioritized: all coding materials and prompts were made openly available, supporting replication and further adaptation of the analytical workflow in other settings. Finally, the black-box nature of LLMs limits insight into how specific outputs were generated. To mitigate this, 2 experts manually reviewed the most positive and most negative-scoring N-grams to assess whether the assigned sentiment scores aligned with their contextual meaning.

Several limitations should also be acknowledged. First, individual narrative descriptions were often brief and fragmented with an average length of 10.5 (SD 5.80) words. This limited the semantic richness available for LLM-based analysis and may have affected the quality of all outputs. However, analyzing the entire corpus at scale partly offset this limitation, allowing recurring patterns to emerge, despite the brevity of individual notes. Second, the data were collected and analyzed in Dutch nursing home settings, which may limit the generalizability and transferability of the findings to other long-term care contexts and languages. However, analyzing the narratives in their original language avoided potential loss of meaning that could have occurred through translation before analysis. Third, although the final prompts were developed through an iterative refinement process, each final prompt was executed only once for the final analysis. LLM-based outputs may vary across repeated executions and model versions, even when prompts and input data remain unchanged [50,96]. We documented the model, deployment, API version, software package version, analysis period, prompts, and code; however, exact computational reproducibility cannot be guaranteed. Fourth, although the expert panel evaluated the sentiment scores of the top-scoring N-grams, the analysis relied on a general-language lexicon rather than a dementia-specific lexicon. Consequently, some context-dependent terms may have been mis-scored and, if they did not appear among the reviewed N-grams, remained underutilized. Finally, the expert review assessed whether selected outputs were coherent, plausible, and meaningful in relation to long-term care practice but did not perform formal validation or allow for the calculation of agreement metrics. Because QoL in dementia is subjective and lacks a single objective gold standard for validation, the findings should be interpreted as exploratory QoL-relevant patterns rather than validated classifications or measurements of QoL [66].

Future Research

Future research should build on these findings in several directions. First, larger and more comprehensive datasets, including greater and richer volumes of observational narrative descriptions, are needed to strengthen the robustness and generalizability of the analysis. Second, a dementia-specific lexicon should be developed and validated to ensure that terms are weighted appropriately in this context. Third, prospective validation studies are required to examine whether LLM-based themes and sentiment patterns correspond to established QoL instrument scores. Fourth, research should explore the real-time integration of LLM-based analysis into observation tools such as MEDLO, evaluating usability, interpretability, and workflow alignment through co-design with care professionals.

These findings should be interpreted as exploratory. Although the current analysis suggests that LLM-based analyses may help structure observational narratives into QoL-relevant patterns, the approach is not yet validated against QoL measures. Before such outputs can inform person-centered care, further research is needed to examine whether LLM-generated outputs correspond to established QoL measures if validated and integrated into care workflows.

Conclusion

LLM-based analyses predominantly identified passive activities in little variation in indoor settings; however, people living with dementia were often described as having a positive affect. These findings suggest that LLM-based analyses can structure observational narratives into interpretable patterns of activity, emotional expression, and social interaction. Further validation is required before such outputs can support multidisciplinary reflection and person-centered dementia care.

Acknowledgments

The authors thank the researchers from the original Green Care Farm project for providing the narrative observation data and for conducting the expert review in this study. Additional support from Alzheimer Nederland is gratefully acknowledged. During preparation and revision of this manuscript, the main author (DS) used ChatGPT (OpenAI) to support language editing and to assist with drafting responses to reviewer comments. This use was limited to writing support. ChatGPT was not used to generate references, make autonomous scientific decisions, or interpret the study findings. The analysis of the observational narrative data was conducted only through the predefined natural language processing pipeline described in the Methods section. All AI-assisted text was critically reviewed, edited, and approved by the authors. The authors take full responsibility for the content of the manuscript.

Funding

This research was funded by the Dutch Research Council (NWO) through the QoLEAD (Quality of Life by use of Enabling AI in Dementia) Project (project number KICH1.GZ02.20.008).

Data Availability

The model and prompts can be found on GitHub.

Authors' Contributions

Conceptualization: DS, MCS, HC, MdV, HV, SA

Data curation: DS

Formal analysis: DS

Funding acquisition: MdV

Investigation: DS, MCS, RFvdW

Methodology: DS, MCS, HC, MvD, HV, SA

Project administration: DS, MdV

Software: RFvdW, MS

Supervision: MvD, HV, MCS

Validation: DS, MCS, HC, SA

Visualization: DS

Writing – original draft: DS

Writing – review and editing: DS, MCS, HC, MdV, HV, SA

Conflicts of Interest

None declared.

  1. Sachdev PS, Blacker D, Blazer DG, et al. Classifying neurocognitive disorders: the DSM-5 approach. Nat Rev Neurol. Nov 2014;10(11):634-642. [CrossRef] [Medline]
  2. Banerjee S, Samsi K, Petrie CD, et al. What do we know about quality of life in dementia? A review of the emerging evidence on the predictive and explanatory value of disease specific measures of health related quality of life in people with dementia. Int J Geriatr Psychiatry. Jan 2009;24(1):15-24. [CrossRef] [Medline]
  3. Lawton MP. A multidimensional view of quality of life in frail elders. In: Birren JE, Rowe JC, Deutchman DE, Lubben JE, editors. The Concept and Measurement of Quality of Life in the Frail Elderly. Elsevier; 1991:3-27. [CrossRef]
  4. Cerejeira J, Lagarto L, Mukaetova-Ladinska EB. Behavioral and psychological symptoms of dementia. Front Neurol. 2012;3:73. [CrossRef] [Medline]
  5. Dröes RM, Boelens-Van Der Knoop EC, Bos J, et al. Quality of life in dementia in perspective: an explorative study of variations in opinions among people with dementia and their professional caregivers, and in literature. Dementia. 2006;5(4):533-558. [CrossRef]
  6. Ettema TP, Dröes RM, de Lange J, Ooms ME, Mellenbergh GJ, Ribbe MW. The concept of quality of life in dementia in the different stages of the disease. Int Psychogeriatr. Sep 2005;17(3):353-370. [CrossRef] [Medline]
  7. Schmüdderich K, Holle D, Ströbel A, Holle B, Palm R. Relationship between the severity of agitation and quality of life in residents with dementia living in German nursing homes - a secondary data analysis. BMC Psychiatry. Apr 13, 2021;21(1):191. [CrossRef] [Medline]
  8. Banerjee S, Smith SC, Lamping DL, et al. Quality of life in dementia: more than just cognition. An analysis of associations with quality of life in dementia. J Neurol Neurosurg Psychiatry. Feb 1, 2006;77(2):146-148. [CrossRef] [Medline]
  9. Hoe J, Hancock G, Livingston G, Orrell M. Quality of life of people with dementia in residential care homes. Br J Psychiatry. May 2006;188(5):460-464. [CrossRef] [Medline]
  10. Scherrer Júnior G, Okuno MFP, de Oliveira LM, et al. Quality of life of institutionalized aged with and without symptoms of depression. Rev Bras Enferm. Nov 2019;72(suppl 2):127-133. [CrossRef] [Medline]
  11. Banovic S, Zunic LJ, Sinanovic O. Communication difficulties as a result of dementia. Mater Sociomed. Oct 2018;30(3):221-224. [CrossRef] [Medline]
  12. Warren A. Behavioral and psychological symptoms of dementia as a means of communication: considerations for reducing stigma and promoting person-centered care. Front Psychol. 2022;13:875246. [CrossRef] [Medline]
  13. Sanghera S, Walther A, Peters TJ, Coast J. Challenges in using recommended quality of life measures to assess fluctuating health: a think-aloud study to understand how recall and timing of assessment influence patient responses. Patient. Jul 2022;15(4):445-457. [CrossRef] [Medline]
  14. Beerens HC. Adding life to years: quality of life of people with dementia receiving long-term care [Doctoral Dissertation]. Maastricht University; 2016. URL: https:/​/cris.​maastrichtuniversity.nl/​en/​publications/​adding-life-to-years-quality-of-life-of-people-with-dementia-rece/​ [Accessed 2026-08-27]
  15. de Boer B, Beerens HC, Zwakhalen SMG, Tan FES, Hamers JPH, Verbeek H. Daily lives of residents with dementia in nursing homes: development of the Maastricht electronic daily life observation tool. Int Psychogeriatr. Aug 2016;28(8):1333-1343. [CrossRef] [Medline]
  16. Nikhil R, Tikoo N, Kurle S, Pisupati HS, Prasad GR. A survey on text mining and sentiment analysis for unstructured web data. J Emerg Technol Innov Res. 2015;2(4):1292-1296. URL: https://www.jetir.org/view?paper=JETIR1504084 [Accessed 2026-08-27]
  17. Adnan K, Akbar R, Khor SW, Ali ABA. Role and challenges of unstructured big data in healthcare. In: Sharma N, Chakrabarti A, Balas VE, editors. Data Management, Analytics and Innovation: Proceedings of ICDMAI 2019, Volume 1. Springer; 2020:301-323. [CrossRef]
  18. Tayefi M, Ngo P, Chomutare T, et al. Challenges and opportunities beyond structured data in analysis of electronic health records. WIREs Comput Stat. Nov 2021;13(6):e1549. [CrossRef]
  19. Hacking C, Verbeek H, Hamers JPH, Aarts S. Comparing text mining and manual coding methods: analysing interview data on quality of care in long-term care for older adults. PLoS ONE. 2023;18(11):e0292578. [CrossRef] [Medline]
  20. Razzaqe MA, Basak T. Text mining in unstructured text: techniques, methods and analysis. World Sci News An Int Sci J. 2022;174:76-92. URL: https:/​/www.​researchgate.net/​profile/​Tapati-Basak/​publication/​375583862_Text_mining_in_unstructured_text_techniques_methods_and_analysis/​links/​6550735c3fa26f66f4f47eca/​Text-mining-in-unstructured-text-techniques-methods-and-analysis.​pdf [Accessed 2026-08-27]
  21. Young T, Hazarika D, Poria S, Cambria E. Recent trends in deep learning based natural language processing. IEEE Comput Intell Mag. 2018;13(3):55-75. [CrossRef]
  22. Brown T, Mann B, Ryder N, et al. Language models are few-shot learners. Presented at: NIPS’20: The 34th International Conference on Neural Information Processing Systems; Dec 6-12, 2020. URL: https://dl.acm.org/doi/abs/10.5555/3495724.3495883 [Accessed 2026-08-27]
  23. Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Presented at: NIPS’17: The 31st International Conference on Neural Information Processing Systems; Dec 4-9, 2017. URL: https://dl.acm.org/doi/10.5555/3295222.3295349 [Accessed 2026-08-27]
  24. Wieland-Jorna Y, van Kooten D, Verheij RA, de Man Y, Francke AL, Oosterveld-Vlug MG. Natural language processing systems for extracting information from electronic health records about activities of daily living. A systematic review. JAMIA Open. Jul 2024;7(2):ooae044. [CrossRef] [Medline]
  25. Feizollah A, Lin CY, O’Malley L, Thompson W, Listl S, Byrne M. The use of natural language processing to interpret unstructured patient feedback on health services: scoping review. J Med Internet Res. Aug 14, 2025;27:e72853. [CrossRef] [Medline]
  26. Lázaro E, Moscardó V. Qualitative health-related quality of life and natural language processing: characteristics, implications, and challenges. Healthcare (Basel). Oct 8, 2024;12(19):2008. [CrossRef] [Medline]
  27. Rickman S, Fernandez JL, Malley J. Understanding patterns of loneliness in older long-term care users using natural language processing with free text case notes. PLoS One. 2025;20(4):e0319745. [CrossRef] [Medline]
  28. Chaichulee S, Promchai C, Kaewkomon T, Kongkamol C, Ingviya T, Sangsupawanich P. Multi-label classification of symptom terms from free-text bilingual adverse drug reaction reports using natural language processing. PLoS ONE. 2022;17(8):e0270595. [CrossRef] [Medline]
  29. Blinov P, Avetisian M, Kokh V, Umerenkov D, Tuzhilin A. Predicting clinical diagnosis from patients electronic health records using BERT-based neural networks. In: Michalowski M, Moskovitch R, editors. Artificial Intelligence in Medicine: 18th International Conference on Artificial Intelligence in Medicine, AIME 2020, Minneapolis, MN, USA, August 25–28, 2020, Proceedings. Springer; 2020:111-121. [CrossRef]
  30. Li Y, Rao S, Solares JRA, et al. BEHRT: transformer for electronic health records. Sci Rep. 2020;10(1):7155. [CrossRef]
  31. Guevara M, Chen S, Thomas S, et al. Large language models to identify social determinants of health in electronic health records. NPJ Digit Med. Jan 11, 2024;7(1):6. [CrossRef] [Medline]
  32. Zhu VJ, Lenert LA, Bunnell BE, Obeid JS, Jefferson M, Halbert CH. Automatically identifying social isolation from clinical narratives for patients with prostate cancer. BMC Med Inform Decis Mak. Mar 14, 2019;19(1):43. [CrossRef] [Medline]
  33. Paek H, Fortinsky RH, Lee K, et al. Real-world insights into dementia diagnosis trajectory and clinical practice patterns unveiled by natural language processing: development and usability study. JMIR Aging. Feb 25, 2025;8:e65221. [CrossRef] [Medline]
  34. Ryvicker M, Barrón Y, Song J, et al. Using natural language processing to identify home health care patients at risk for diagnosis of Alzheimer’s disease and related dementias. J Appl Gerontol. Oct 2024;43(10):1461-1472. [CrossRef] [Medline]
  35. Zhang H, Vithanage D, Song T, Deng C, Yu P. Leveraging retrieval augmented generation-driven large language models to extract dementia agitation symptoms and triggers from free-text nursing notes. Stud Health Technol Inform. Aug 7, 2025;329:799-803. [CrossRef] [Medline]
  36. Rosteius K, de Boer B, Steinmann G, Verbeek H. What can other dementia care settings learn from Green Care Farms? An exploration of their functions and forms. Innov Aging. Dec 21, 2023;7(Supplement_1):233-234. [CrossRef]
  37. Verbeek H, Zwakhalen SMG, Schols JMGA, Kempen GIJM, Hamers JPH. The living lab in ageing and long-term care: a sustainable model for translational research improving quality of life, quality of care and quality of work. J Nutr Health Aging. 2020;24(1):43-47. [CrossRef] [Medline]
  38. World Medical Association. World Medical Association Declaration of Helsinki: ethical principles for medical research involving human participants. JAMA. Jan 7, 2025;333(1):71-74. [CrossRef] [Medline]
  39. Hofmann M, Chisholm A. Text Mining and Visualization: Case Studies Using Open-Source Tools. CRC Press; 2016. ISBN: 9781482237580
  40. Han J, Kamber M, Pei J. Data Mining: Concepts and Techniques. Morgan kaufmann; 2022. ISBN: 978-0-12-381479-1
  41. van Rossum G. Python reference manual. Centrum voor Wiskunde en Informatica (CWI); 1995. URL: https://ir.cwi.nl/pub/5008/05008D.pdf [Accessed 2026-08-27]
  42. QoLEAD llms-based secure and reliable workflow. GitHub. 2024. URL: https://github.com/HR-DataLab-Healthcare/RESEARCH_SUPPORT/tree/main/PROJECTS/QoLEAD [Accessed 2026-08-27]
  43. Manning CD, Raghavan P, Schütze H. Introduction to Information Retrieval. Cambridge university press; 2008. URL: https://www.cambridge.org/core/product/identifier/9780511809071/type/book [Accessed 2026-08-27]
  44. Camacho-Collados J, Pilehvar MT. On the role of text preprocessing in neural network architectures: an evaluation study on text categorization and sentiment analysis. In: Linzen T, Chrupała G, Alishahi A, editors. Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. Association for Computational Linguistics; 2018:40-46. [CrossRef]
  45. Microsoft. 2026. URL: https://www.microsoft.com/en-in [Accessed 2026-08-27]
  46. OpenAI. 2025. URL: https://openai.com/ [Accessed 2026-08-27]
  47. Sánchez AG. Azure OpenAI Service for Cloud Native Applications: Designing, Planning, and Implementing Generative AI Solutions. O’Reilly Media, Inc; 2024. URL: https://www.oreilly.com/library/view/azure-openai-service/9781098154981/ [Accessed 2026-08-27]
  48. McKinney W. Pandas: a foundational python library for data analysis and statistics. Presented at: Python for High Performance and Scientific Computing (PyHPC 2011); Nov 18, 2011. URL: https:/​/www.​researchgate.net/​publication/​265194455_pandas_a_Foundational_Python_Library_for_Data_Analysis_and_Statistics [Accessed 2026-08-31]
  49. PyPDF2 301. 2026. URL: https://pypi.org/project/PyPDF2/ [Accessed 2026-08-31]
  50. Learn how to use reproducible output (preview) (classic). Microsoft. 2026. URL: https:/​/learn.​microsoft.com/​en-us/​azure/​foundry-classic/​openai/​how-to/​reproducible-output?tabs=python [Accessed 2026-08-31]
  51. Structured outputs. Microsoft. 2026. URL: https:/​/learn.​microsoft.com/​en-us/​azure/​foundry/​openai/​how-to/​structured-outputs?source=recommendations&tabs=python-secure%2Cdotnet-entra-id%2Cjavascript-secure&pivots=programming-language-python [Accessed 2026-08-27]
  52. ChatCompletionsOptions.temperature property. Microsoft Learn. 2026. URL: https:/​/learn.​microsoft.com/​en-us/​dotnet/​api/​azure.​ai.​inference.​chatcompletionsoptions.​temperature?view=azure-dotnet-preview [Accessed 2026-08-27]
  53. Madaan A, Tandon N, Gupta P, et al. Self-refine: iterative refinement with self-feedback. Presented at: Advances in Neural Information Processing Systems 36; Dec 10-16, 2023. [CrossRef]
  54. Manning C, Schutze H. Foundations of Statistical Natural Language Processing. MIT press; 1999. ISBN: 9780262133609
  55. Hotho A, Nürnberger A, Paaß G. A brief survey of text mining. J Lang Technol Comput Linguistics. 2005;20(1):19-62. [CrossRef]
  56. Reagan AJ, Danforth CM, Tivnan B, Williams JR, Dodds PS. Sentiment analysis methods for understanding large-scale texts: a case for using continuum-scored words and word shift graphs. EPJ Data Sci. Dec 2017;6(1):28. [CrossRef]
  57. Bock HH. Clustering methods: a history of k-means algorithms. In: Brito P, Cucumel G, Bertrand P, de Carvalho F, editors. Selected Contributions in Data Analysis and Classification. Springer; 2007:161-172. [CrossRef]
  58. Kherwa P, Bansal P. Topic modeling: a comprehensive review. EAI Endorsed Trans Scalable Inf Syst. 2019;7(24):e2. [CrossRef]
  59. Mu Y, Dong C, Bontcheva K, Song X. Large language models offer an alternative to the traditional approach of topic modelling. In: Calzolari N, Kan MY, Hoste V, Lenci A, Sakti S, Xue N, editors. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). ELRA and ICCL; 2024:10160-10171. URL: https://aclanthology.org/2024.lrec-main.887/ [Accessed 2026-08-27]
  60. Kozlowski D, Pradier C, Benz P. Generative AI for automatic topic labelling. Can J Inf Libr Sci. 2026;49(2):26-34. [CrossRef]
  61. Gale NK, Heath G, Cameron E, Rashid S, Redwood S. Using the framework method for the analysis of qualitative data in multi-disciplinary health research. BMC Med Res Methodol. Sep 18, 2013;13:117. [CrossRef] [Medline]
  62. Frissen L, Aarts S, Rosteius K, de Boer B, Gabrio A, Verbeek H. The influence of social interactions on mood in residents with dementia in Green Care Farms: an observational study using ecological momentary assessments. Int Psychogeriatr. Sep 2025;37(5):100091. [CrossRef] [Medline]
  63. Brooker D, Latham I. Person-Centred Dementia Care: Making Services Better with the VIPS Framework. Jessica Kingsley Publishers; 2015. ISBN: 9781849056663
  64. van Wijngaarden E, van der Wedden H, Henning Z, Komen R, The AM. Entangled in uncertainty: the experience of living with dementia from the perspective of family caregivers. PLoS ONE. 2018;13(6):e0198034. [CrossRef] [Medline]
  65. Hofbauer LM, Rodriguez FS. Psychosocial wellbeing of people with dementia: systematic review and construct analysis. Acta Neuropsychiatr. Jun 23, 2025;37:e71. [CrossRef] [Medline]
  66. Jonker C, Gerritsen DL, Bosboom PR, Van Der Steen JT. A model for quality of life measures in patients with dementia: Lawton’s next step. Dement Geriatr Cogn Disord. 2004;18(2):159-164. [CrossRef] [Medline]
  67. Kitwood T. The experience of dementia. Aging Ment Health. Feb 1997;1(1):13-22. [CrossRef]
  68. Young JC, Arthur R, Williams HTP. CIDER: context-sensitive polarity measurement for short-form text. PLoS ONE. 2024;19(4):e0299490. [CrossRef] [Medline]
  69. Loughran T, Mcdonald B. When is a liability not a liability? Textual analysis, dictionaries, and 10‐Ks. J Finance. Feb 2011;66(1):35-65. [CrossRef]
  70. Liu J, Capurro D, Nguyen A, Verspoor K. “Note Bloat” impacts deep learning-based NLP models for clinical prediction tasks. J Biomed Inform. Sep 2022;133:104149. [CrossRef] [Medline]
  71. Sohn S, Wang Y, Wi CI, et al. Clinical documentation variations and NLP system portability: a case study in asthma birth cohorts across institutions. J Am Med Inform Assoc. Mar 1, 2018;25(3):353-359. [CrossRef] [Medline]
  72. Scharp D, Hobensack M, Davoudi A, Topaz M. Natural language processing applied to clinical documentation in post-acute care settings: a scoping review. J Am Med Dir Assoc. Jan 2024;25(1):69-83. [CrossRef] [Medline]
  73. Kaufmann EG, Engel SA. Dementia and well-being: a conceptual framework based on Tom Kitwood’s model of needs. Dementia (London). Jul 2016;15(4):774-788. [CrossRef] [Medline]
  74. Han A, Radel J, McDowd JM, Sabata D. Perspectives of people with dementia about meaningful activities: a synthesis. Am J Alzheimers Dis Other Demen. Mar 2016;31(2):115-123. [CrossRef] [Medline]
  75. Rosteius K, de Boer B, Steinmann G, Verbeek H. Fostering an active daily life: an ethnographic study unravelling the mechanisms of Green Care Farms as innovative long-term care environment for people with dementia. Int Psychogeriatr. Mar 2025;37(2):100017. [CrossRef] [Medline]
  76. Kitwood T. Towards a theory of dementia care: the interpersonal process. Ageing Soc. Mar 1993;13(1):51-67. [CrossRef]
  77. Edvardsson D, Winblad B, Sandman PO. Person-centred care of people with severe Alzheimer’s disease: current status and ways forward. Lancet Neurol. Apr 2008;7(4):362-367. [CrossRef] [Medline]
  78. Kane RA. Long-term care and a good quality of life: bringing them closer together. Gerontologist. Jun 2001;41(3):293-304. [CrossRef] [Medline]
  79. Beerens HC, de Boer B, Zwakhalen SMG, et al. The association between aspects of daily life and quality of life of people with dementia living in long-term care facilities: a momentary assessment study. Int Psychogeriatr. Aug 2016;28(8):1323-1331. [CrossRef] [Medline]
  80. Lawton MP. Assessing quality of life in Alzheimer disease research. Alzheimer Dis Assoc Disord. 1997;11 Suppl 6:91-99. [Medline]
  81. Ettema TP, Dröes RM, de Lange J, Mellenbergh GJ, Ribbe MW. QUALIDEM: development and evaluation of a dementia specific quality of life instrument--validation. Int J Geriatr Psychiatry. May 2007;22(5):424-430. [CrossRef] [Medline]
  82. Smith SC, Lamping DL, Banerjee S, et al. Measurement of health-related quality of life for people with dementia: development of a new instrument (DEMQOL) and an evaluation of current methodology. Health Technol Assess. Mar 2005;9(10):1-93. [CrossRef] [Medline]
  83. Habbal S, Mian M, Imam M, Tahiri J, Amor A, Reddy PH. Harnessing artificial intelligence for transforming dementia care: innovations in early detection and treatment. Brain Organoid Syst Neurosci J. Dec 2025;3:122-133. [CrossRef]
  84. Jerfy A, Selden O, Balkrishnan R. The growing impact of natural language processing in healthcare and public health. Inquiry. 2024;61:469580241290095. [CrossRef] [Medline]
  85. Shankar R, Bundele A, Mukhopadhyay A. Natural language processing of electronic health records for early detection of cognitive decline: a systematic review. NPJ Digit Med. Mar 1, 2025;8(1):133. [CrossRef] [Medline]
  86. Hendriks A, Hacking C, Verbeek H, Aarts S. Data science techniques to gain novel insights into quality of care: a scoping review of long-term care for older adults. Explor Digit Health Technol. 2024;2(2):67-85. [CrossRef]
  87. Shakeri A, Farmanbar M. Natural language processing in Alzheimer’s disease research: systematic review of methods, data, and efficacy. Alzheimers Dement (Amst). Jan 2025;17(1):e70082. [CrossRef] [Medline]
  88. Mitchell G, Agnelli J. Person-centred care for people with dementia: Kitwood reconsidered. Nurs Stand. Oct 14, 2015;30(7):46-50. [CrossRef] [Medline]
  89. Brooker D. What is person-centred care in dementia? Rev Clin Gerontol. Aug 2003;13(3):215-222. [CrossRef]
  90. Sokol K, Fackler J, Vogt JE. Artificial intelligence should genuinely support clinical reasoning and decision making to bridge the translational gap. NPJ Digit Med. Jun 10, 2025;8(1):345. [CrossRef] [Medline]
  91. Sezgin E. Artificial intelligence in healthcare: complementing, not replacing, doctors and healthcare providers. Digit Health. 2023;9:20552076231186520. [CrossRef] [Medline]
  92. Lukkien DRM, Ipakchian Askari S, Stolwijk NE, et al. Making co-design more responsible: case study on the development of an AI-based decision support system in dementia care. JMIR Hum Factors. Jul 31, 2024;11:e55961. [CrossRef] [Medline]
  93. Zicari RV, Ahmed S, Amann J, et al. Co-design of a trustworthy AI system in healthcare: deep learning based skin lesion classifier. Front Hum Dyn. 2021;3:2021. [CrossRef]
  94. Silvola S, Restelli U, Bonfanti M, Croce D. Co-design as enabling factor for patient-centred healthcare: a bibliometric literature review. Clinicoecon Outcomes Res. 2023;15:333-347. [CrossRef] [Medline]
  95. Steijger D, Christie H, Aarts S, IJselsteijn W, Verbeek H, de Vugt M. Use of artificial intelligence to support quality of life of people with dementia: a scoping review. Ageing Res Rev. Jun 2025;108:102741. [CrossRef] [Medline]
  96. Herrera-Poyatos D, Peláez-González C, Zuheros C, et al. An overview of model uncertainty and variability in LLM-based sentiment analysis: challenges, mitigation strategies, and the role of explainability. Front Artif Intell. 2025;8:1609097. [CrossRef] [Medline]


LLM: large language model
MEDLO: Maastricht Electronic Daily Life Observation
NLP: natural language processing
QoL: quality of life


Edited by Ivan Steenstra; submitted 22.Jan.2026; peer-reviewed by Miloud Chakit, Sundeep Venkatesan; final revised version received 17.Jul.2026; accepted 17.Jul.2026; published 15.Sep.2026.

Copyright

© Dirk Steijger, Mark C Scheper, Robert Frans van der Willigen, Hannah Christie, Marjolein E de Vugt, Hilde Verbeek, Sil Aarts. Originally published in the Journal of Medical Internet Research (https://www.jmir.org), 15.Sep.2026.

This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in the Journal of Medical Internet Research (ISSN 1438-8871), is properly cited. The complete bibliographic information, a link to the original publication on https://www.jmir.org/, as well as this copyright and license information must be included.